Papers with natural language processing techniques

20 papers
BiomedCurator: Data Curation for Biomedical Literature (2022.aacl-demo)

Copied to clipboard

Challenge: BiomedCurator uses state-of-the-art natural language processing techniques to extract structured data from scientific articles.
Approach: They propose a web application that extracts structured data from PubMed and ClinicalTrials.gov . the application uses a combination of natural language processing techniques and a pattern-based extraction approach .
Outcome: The proposed system extracts the structured data from PubMed and ClinicalTrials.gov datasets.
Supporting Complaints Investigation for Nursing and Midwifery Regulatory Agencies (2021.acl-demo)

Copied to clipboard

Challenge: Fig. 1 illustrates the major components and workflow of our proposed system to improve the efficiency of complaints investigation for nursing and midwifery regulators.
Approach: They propose a decision support system that uses machine learning and natural language processing techniques to process complaints and predict their risk level.
Outcome: The proposed system uses state-of-the-art machine learning and natural language processing techniques to process complaints and predict risk levels.
Joint Dialogue Topic Segmentation and Categorization: A Case Study on Clinical Spoken Conversations (2023.emnlp-industry)

Copied to clipboard

Challenge: Utilizing natural language processing in clinical conversations is effective to improve the efficiency of workflows for medical staff and patients.
Approach: They propose a model for dialogue segmentation and topic categorization that integrates natural language processing techniques into a joint model.
Outcome: The proposed model improves on follow-up calls for diabetes management and reduces computational complexity and cost.
Generating Image Captions in Arabic using Root-Word Based Recurrent Neural Networks and Deep Neural Networks (N18-4)

Copied to clipboard

Challenge: Existing studies on image caption generation in English focus on Western languages, ignoring Semitic and Middle-Eastern languages like Arabic, Hebrew, Urdu and Persian.
Approach: They propose to leverage the critical dependency of Arabic to generate Arabic captions using root-word based Recurrent Neural Network and Deep Neural networks.
Outcome: The proposed model outperforms English-Arabic translated captions on a dataset from newspapers in the Middle East.
Visualizing Trends of Key Roles in News Articles (D19-3)

Copied to clipboard

Challenge: a demonstration system visualizes news trend of key roles based on natural language processing techniques . semantic role labelling and word embeddings can help users understand news topics .
Approach: They propose a system that visualizes the news trend of key roles based on natural language processing techniques.
Outcome: The proposed system analyzes the news trend of key roles using semantic role labelling . it also analyzes how similarities between key roles and news topics change over time .
DISPUTool 3.0: Fallacy Detection and Repairing in Argumentative Political Debates (2025.acl-demo)

Copied to clipboard

Challenge: DISPUTool 3.0 is a web-based application for identifying and fixing fallacious arguments in political debates.
Approach: They propose a web-based application designed to identify and repair fallacious arguments in political debates.
Outcome: The proposed tool is based on the ElecDeb60to20 dataset covering US presidential debates from 1960 to 2020.
A Natural Approach for Synthetic Short-Form Text Analysis (2024.lrec-main)

Copied to clipboard

Challenge: Social media and news sites can be flooded with synthetically generated misinformation via tweets and posts while authentic users can inadvertently spread this text via shares and retweets.
Approach: They propose a method of detecting synthetically generated tweets via a Transformer architecture and incorporate unique style-based features.
Outcome: The proposed method detects synthetically generated tweets using a Transformer architecture and incorporating unique style-based features.
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to enhance zero-shot abilities in image captioning fail with fine-grained datasets.
Approach: They propose a method to enhance captions with additional object-part details using object detector proposals and natural language processing techniques.
Outcome: The proposed method improves performance on fine-grained datasets and improves on existing methods.
Enhancing Air Quality Prediction with Social Media and Natural Language Processing (P19-1)

Copied to clipboard

Challenge: predicting air quality is a major concern for human health, but the changes of air quality conditions are still difficult to monitor.
Approach: They propose to exploit social media and natural language processing techniques to enhance air quality prediction.
Outcome: The proposed approach improves air quality prediction over baseline that does not use social media by 6.9% to 17.7% in macro-F1 scores.
ChartThinker: A Contextual Chain-of-Thought Approach to Optimized Chart Summarization (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for chart summarization lack visual-language matching and reasoning ability.
Approach: They propose a method which synthesizes deep analysis based on chains of thought and strategies of context retrieval to improve the logical coherence and accuracy of the generated summaries.
Outcome: The proposed method outperforms 8 state-of-the-art models over 7 evaluation metrics and can significantly reduce time and cognitive resources required.
Building an Ellipsis-aware Chinese Dependency Treebank for Web Text (L18-1)

Copied to clipboard

Challenge: ellipsis is a common linguistic phenomenon that some words are left out as they are understood from the context, especially in oral utterance.
Approach: They propose to use a Chinese dependency treebank to facilitate the parsing of web text . they propose to restore omissions and reserve contexts in the web text to improve dependency parsers .
Outcome: The proposed framework enables the parsing of web text from online microblogs.
BPM_MT: Enhanced Backchannel Prediction Model using Multi-Task Learning (2021.emnlp-main)

Copied to clipboard

Challenge: Backchannel (BC) is a short and quick reaction signal of a listener to a speaker's utterances.
Approach: They propose a model that utilizes lexical information in utterances to enhance backchannel (BC) prediction.
Outcome: The proposed model showed 14.24% performance improvement compared to baseline in the four BC categories: continuer, understanding, empathic response, and No BC.
Towards Intention Understanding in Suicidal Risk Assessment with Natural Language Processing (2022.findings-emnlp)

Copied to clipboard

Challenge: Suicide is a global problem, with one suicide case for every 100 deaths worldwide . social networking sites are an essential forum for communication and information sharing .
Approach: This paper compares natural language processing to suicidal ideation detection and risk assessment . it urges better intention understanding for reliable suicide risk assessment with computational methods .
Outcome: This paper compares the performance of natural language processing to suicidal ideation detection and risk assessment tasks.
Evaluating Sentence Segmentation in Different Datasets of Neuropsychological Language Tests in Brazilian Portuguese (2020.lrec-1)

Copied to clipboard

Challenge: Using automated analysis of connected speech is a promising direction for diagnosing cognitive impairments.
Approach: They propose to use a novel model to segment impaired speech transcriptions . they propose to include a Linear Chain CRF and a self-attention mechanism .
Outcome: The proposed system performs better than the existing model with three new datasets used to diagnose cognitive impairments.
Towards Processing of the Oral History Interviews and Related Printed Documents (L18-1)

Copied to clipboard

Challenge: a project aims to create an integrated archive of the recordings, scanned documents and photographs from totalitarian regimes in Czechoslovakia . the archive will be accessible online and provide multifaceted search capabilities .
Approach: They propose to use automatic speech recognition and optical character recognition to build an archive of the interviews, scanned documents and photographs.
Outcome: The proposed archive will be accessible online and provide multifaceted search capabilities.
MIND: A Large-scale Dataset for News Recommendation (2020.acl-main)

Copied to clipboard

Challenge: Personalized news recommendation is an important technique for personalized news service.
Approach: They propose to build a large-scale news recommendation dataset from Microsoft News . they demonstrate that news recommendation relies on the quality of news content understanding .
Outcome: The proposed dataset contains 1 million users and more than 160k English news articles, each of which has rich textual content such as title, abstract and body.
Predicting Anti-Asian Hateful Users on Twitter during COVID-19 (2021.findings-emnlp)

Copied to clipboard

Challenge: Xenophobia and polarization have accompanied widespread social media usage in many nations, attracting many researchers.
Approach: They apply natural language processing techniques to characterize Twitter users who began to post anti-Asian hate messages during COVID-19.
Outcome: The results show that it is possible to predict who later posted anti-Asian slurs on Twitter and Reddit.
Why Swear? Analyzing and Inferring the Intentions of Vulgar Expressions (D18-1)

Copied to clipboard

Challenge: Vulgar words are employed in language use for several different functions, including expressing aggression, signaling group identity or the informality of the communication.
Approach: They present a dataset of 7,800 tweets with six categories of vulgarity in which all instances of vulgar words are annotated with one of the six categories.
Outcome: The proposed model can predict the category of a vulgar word based on the immediate context it appears in with 67.4 macro F1 across six classes.
What to Fuse and How to Fuse: Exploring Emotion and Personality Fusion Strategies for Explainable Mental Disorder Detection (2023.findings-acl)

Copied to clipboard

Challenge: Mental health disorders (MHD) are one of the greatest challenges facing our healthcare systems and modern societies in general.
Approach: They integrate and extend the research by conducting extensive experiments with three types of deep learning-based fusion strategies: feature-level fusion, model fusion and task fusion.
Outcome: The proposed techniques show that they can be used to improve mental health detection from textual data.
CTAP for Italian: Integrating Components for the Analysis of Italian into a Multilingual Linguistic Complexity Analysis Tool (2020.lrec-1)

Copied to clipboard

Challenge: Linguistic complexity is a core construct in Second Language Acquisition (SLA) research.
Approach: They present an open source linguistic complexity measurement tool for Italian . they compare it to existing tools for English and germany .
Outcome: The proposed tool is the most comprehensive linguistic complexity measurement tool for italian . it can be used to compare italian texts to multiple other languages in one tool .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations